When Evidence Travels: Reflections from GIN 2026 

By Kinlabel Tetamiyaka Okwen

Conferences can sometimes feel like a strange compression of time. You spend months working through ideas, methods, failed experiments, half-finished diagrams and difficult questions, and then suddenly you have ten or fifteen minutes in a room to explain what all of it means. 

Guidelines International Conference 2026 felt like that for me. I went into the conference with two pieces of work that, on the surface, were quite different. One was about keeping evidence alive after a review is completed. The other was about whether evidence and recommendations can travel from one context to another. By the end of the conference, I realised they were really asking the same question: What happens to evidence after we produce it? 

Does it sit in a paper? Does it become a guideline that is copied elsewhere? Does it change when the context changes? Does someone have to start from scratch five years later? And increasingly, where does AI fit into all of this? Those questions followed me through GIN 2026 in Malaysia, which was the first GIN conference I had been able to attend. It was genuinely an experience, not only because I was presenting work I had spent a lot of time thinking about, but because I was suddenly hearing those ideas challenged, stretched, and interpreted by people who approach evidence and guideline development from very different perspectives. 

Don’t Let the Evidence Die 

My oral presentation, Don’t Let the Evidence Die: From Rapid to Living Reviews, Experiences From the Global South, drew on work developed and written with my co-authors. I presented it alone at GIN 2026, and the work grew from a frustration that I think many evidence synthesis teams will recognise. We put enormous effort into reviews: developing protocols, searching databases, screening hundreds or thousands of records, extracting data and synthesising findings to produce something we hope will support a decision. Then the review is published, but the evidence keeps moving. 

COVID-19 made this problem particularly visible. During the pandemic, evidence could change so quickly that a review risked becoming outdated almost immediately. New studies appeared, recommendations shifted, and decision-makers needed evidence at a speed that traditional review cycles were never designed to accommodate. Most evidence questions do not move at COVID speed, but COVID exposed a much bigger methodological problem: our evidence products are often static, while the evidence itself is not. 

Our work explored what it might look like to design rapid reviews with updating in mind from the beginning. Instead of treating each review as a one-off product, we started thinking about the protocol, screening decisions, extraction structures and search strategies as reusable infrastructure. That does not mean a protocol should remain unchanged forever. Research questions evolve, methods improve and context changes. A reusable protocol is not a frozen protocol. It is a methodological backbone that can be amended, documented and versioned rather than rebuilt from zero every time new evidence appears. 

This is also where AI becomes particularly useful because living evidence systems contain enormous amounts of repetitive work: surveillance, prioritisation, screening, extraction and updating. AI and machine learning can potentially help with some of this work, but the goal is not automation for its own sake. The more important question is whether automation makes the process more efficient, sustainable and reliable while maintaining appropriate human oversight. For me, the more important question is not simply, “Can AI do this?” but “Under what conditions should we trust it to do this?” Where does the model perform well? Where does it fail? Which extraction fields are reliable? How much human-labelled data is required before a screening classifier becomes useful? When should we abandon automation and return to full human review? 

These questions are explored in much more detail in our paper, Don’t Let the Evidence Die: From Rapid to Living Reviews, Experiences From the Global South. 

Kinlabel presenting the paper "Don't Let the Evidence Die" at GIN2026 Conference

Can guidelines travel? 

My second major conversation at GIN was about transferability. Using the NICE Type 2 Diabetes guideline, our workshop explored a deceptively simple question: if a recommendation works somewhere else, can we use it here? 

This was the first time we had used our transferability model to assess a guideline recommendation rather than a research study, and the distinction mattered. A study usually presents a defined intervention, population, comparator and set of outcomes. A recommendation may sit on top of multiple studies, judgements about certainty, values and preferences, implementation considerations and several intervention options. Deciding what should feed into a transferability assessment is therefore not straightforward. 

For the NICE recommendation, we needed input from someone with professional expertise in Type 2 Diabetes to identify the evidence and assumptions most important to it (Thank you Dr. Okwen!).) . This showed us that assessing the transferability of a study is not necessarily the same as assessing a recommendation. Recommendations carry layers of interpretation and judgement that the model must learn to recognise. 

Guidelines are sometimes discussed as though the evidence embedded within them can simply cross borders. But a recommendation may have been developed in a health system with very different resources, workforce capacity, costs, infrastructure and patient priorities. The evidence may travel. The recommendation may not. 

Kinlabel presenting during a workshop on Guidelines for all: Bridging tradition, technology and high value care, NICE Recommendation. 

Our model tries to make this problem explicit through five predictors: relevance, cost, complexity, impact and importance. The question is not only, “Does this intervention work?” It is also: Is it relevant to the receiving context? Can the system afford it? How difficult would it be to implement? What impact might it have locally? How important is the problem or intervention in that setting? 

This is where our Country and City Packs become important. Transferability cannot be assessed meaningfully if “context” is reduced to a few sentences. The packs provide structured information about the health system, governance, workforce, infrastructure, economic conditions, population and policy environment. This information is given to the language model alongside the evidence, helping turn context from an abstract idea into something the model can reason with. 

Around 20 people participated in the workshop, giving us enough space for genuine discussion. We used the NICE guideline as a practical example and demonstrated how our AI-assisted Playground could interpret the evidence and context, classify the predictors and generate an output. 

The feedback was particularly useful. Participants suggested adding short explanations of each predictor to make the Playground easier to understand. Others asked whether guideline developers should see all the model’s outputs rather than only the final decision to “adopt”, “adapt” or “generate new evidence”. 

Kinlabel interacts with participants at the workshop

That question stayed with me because the reasoning may be more valuable than the final score. A guideline developer may need to know that a recommendation is highly relevant but difficult to implement, or that its potential impact is high while affordability remains uncertain. These individual signals show where adaptation is needed. 

Another important question was where the boundary should sit between global reuse and local development. How much of an existing guideline can we responsibly carry into another setting, and when are the contextual differences too great for adaptation to be enough? If an intervention remains relevant, affordable, feasible and likely to have similar effects, much of the evidence may be reusable. But major differences in workforce, infrastructure, costs, population needs or implementation conditions may require additional local evidence or a locally developed recommendation. 

The comparator in PICO is also crucial. We are not assessing an intervention in isolation; we are comparing it with what already exists in the receiving setting. The realistic alternative may be usual care, a lower-cost intervention or no intervention at all. Understanding that comparator helps us determine whether adaptation is possible before concluding that entirely new evidence is required. 

Our greatest limitation was time. I would have liked to hear more about whether participants believed NICE recommendations could genuinely travel into their contexts. Transferability cannot be tested through a model alone. It must also be tested against how guideline developers, policymakers and evidence users reason about context. 

I explored similar questions on a panel focused on Cameroon: Do international recommendations address the decisions the country needs to make? What can be reused, and what requires local adaptation or development? The discussion reinforced that the main barriers are not always a lack of evidence. They may be missing information about workforce, infrastructure, affordability, implementation capacity and available alternatives. 

That panel made the purpose of the transferability model feel particularly concrete. The model is not intended to automate away human judgement. It is meant to make that judgement more explicit, structured and transparent by showing which contextual factors support transfer, which create uncertainty and which make adaptation necessary. 

What I took away from GIN 2026 was not a neat answer about whether evidence can stay alive or guidelines can travel. It was a clearer understanding of how closely connected these problems are. Evidence changes over time, but context changes too. AI can help us manage some of that complexity, but only if we test its limits, make its reasoning visible and keep human judgement firmly within the process. 

We spend a lot of time asking how to produce better evidence. Increasingly, I think we must also ask how evidence lives after it is produced: how it is updated, interpreted, challenged, adapted and eventually used somewhere very different from where it began. GIN 2026 gave me an opportunity to present some of our answers, but more importantly, it gave me better questions to take back into the work.     


Admin eBASE

eBASE Administrator, responsible for ensuring our digital presence is effective and up-to-date. Manages the content management system, oversees website updates, and analyzes key metrics to improve user experience.

Plays a crucial role in showcasing our mission and impact to a global audience, ensuring our online platform is a reliable and accessible resource for our community, partners, and stakeholders.